Papers with digital devices

6 papers
Luring as a Proxy: Evaluating Corpus Transferability for Cybergrooming Detection (2026.acl-short)

Copied to clipboard

Challenge: Prior research has noted that cybergrooming often involves the use of luring communication strategies to manipulate minors.
Approach: They examine the potential transferability of corpora from luring contexts for cybergrooming detection.
Outcome: The proposed model can be generalized across domains and performs better on corpora with salient toxicity and distinctive stylistic features.
INREACT: An Inspire-Then-Reinforce Training Framework For Multimodal GUI Agent (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing multimodal large language models struggle with precise localization of small elements.
Approach: They propose a multimodal GUI agent framework that unifies observing, thinking, and acting for precise and interpretable decision-making.
Outcome: The proposed framework unifies observing, thinking, and acting for precise and interpretable decision-making.
SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents (2024.acl-long)

Copied to clipboard

Challenge: Existing GUI agents interact with the environment through extracted structured data, which can be notably lengthy (e.g., HTML) and occasionally inaccessible (e-book).
Approach: They propose to enhance SeeClick with GUI grounding pre-training and devise a method to automate curation of GUI ground data.
Outcome: The proposed agent improves ScreenSpot, the first realistic GUI grounding benchmark that encompasses mobile, desktop, and web environments.
Read Anywhere Pointed: Layout-aware GUI Screen Reading with Tree-of-Lens Grounding (2024.emnlp-main)

Copied to clipboard

Challenge: Existing models for GUI understanding ignore a key GUI-referring task: screen reading based on user-indicated points.
Approach: They propose a Tree-of-Lens agent that constructs a Hierarchical Layout Tree based on user input points and a GUI screenshot.
Outcome: The proposed agent can interpret the Screen Point-and-Read task on mobile, web, and operating systems.
Researching Less-Resourced Languages – the DigiSami Corpus (L18-1)

Copied to clipboard

Challenge: DigiSami project aims to support research on endangered languages . it uses spoken corpus and speech technology for the Fenno-Ugric language North Sami .
Approach: They describe the DigiSami project and its research results for the Fenno-Ugric language North Sami . they discuss ethical and privacy issues related to data collection for less-resourced languages and indigenous communities .
Outcome: The DigiSami project focuses on spoken corpus collection and speech technology for the Fenno-Ugric language North Sami.
UI-E2I-Synth: Advancing GUI Grounding with Large-Scale Instruction Synthesis (2025.findings-acl)

Copied to clipboard

Challenge: Graphical User Interface (GUI) agents that utilize human-like vision perception capabilities are gaining a wider applicability compared to GUI metadata-based approaches.
Approach: They propose a large-scale data synthesis pipeline for generating varying complex instruction datasets using GPT-4o instead of human annotators.
Outcome: The proposed model achieves superior performance in GUI instruction grounding, demonstrating the advancements of proposed data synthesis pipeline.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations